Goto

Collaborating Authors

 video processing


StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding

arXiv.org Artificial Intelligence

Multimodal large language models (MLLMs) have made significant progress in visual-language reasoning, but their ability to efficiently handle long videos remains limited. Despite recent advances in long-context MLLMs, storing and attending to the key-value (KV) cache for long visual contexts incurs substantial memory and computational overhead. Existing visual compression methods require either encoding the entire visual context before compression or having access to the questions in advance, which is impractical for long video understanding and multi-turn conversational settings. In this work, we propose StreamMem, a query-agnostic KV cache memory mechanism for streaming video understanding. Specifically, StreamMem encodes new video frames in a streaming manner, compressing the KV cache using attention scores between visual tokens and generic query tokens, while maintaining a fixed-size KV memory to enable efficient question answering (QA) in memory-constrained, long-video scenarios. Evaluation on three long video understanding and two streaming video question answering benchmarks shows that StreamMem achieves state-of-the-art performance in query-agnostic KV cache compression and is competitive with query-aware compression approaches.


A Survey on Video Analytics in Cloud-Edge-Terminal Collaborative Systems

arXiv.org Artificial Intelligence

The explosive growth of video data has driven the development of distributed video analytics in cloud-edge-terminal collaborative (CETC) systems, enabling efficient video processing, real-time inference, and privacy-preserving analysis. Among multiple advantages, CETC systems can distribute video processing tasks and enable adaptive analytics across cloud, edge, and terminal devices, leading to breakthroughs in video surveillance, autonomous driving, and smart cities. In this survey, we first analyze fundamental architectural components, including hierarchical, distributed, and hybrid frameworks, alongside edge computing platforms and resource management mechanisms. Building upon these foundations, edge-centric approaches emphasize on-device processing, edge-assisted offloading, and edge intelligence, while cloud-centric methods leverage powerful computational capabilities for complex video understanding and model training. Our investigation also covers hybrid video analytics incorporating adaptive task offloading and resource-aware scheduling techniques that optimize performance across the entire system. Beyond conventional approaches, recent advances in large language models and multimodal integration reveal both opportunities and challenges in platform scalability, data protection, and system reliability. Future directions also encompass explainable systems, efficient processing mechanisms, and advanced video analytics, offering valuable insights for researchers and practitioners in this dynamic field.


Demo: RhythmEdge: Enabling Contactless Heart Rate Estimation on the Edge

arXiv.org Artificial Intelligence

Our RhythmEdge system is portable and easily deployable for reliable HR estimation in moderately controlled indoor or outdoor environments. RhythmEdge measures HR via detecting changes in blood volume from facial videos (Remote Photoplethysmography; rPPG) and provides instant assessment using off-the-shelf commercially available resource-constrained edge platforms and video cameras. We demonstrate the scalability, flexibility, and compatibility of the RhythmEdge by deploying it on three resource-constrained platforms of differing architectures (NVIDIA Jetson Nano, Google Coral Development Board, Raspberry Pi) and three heterogeneous cameras of differing sensitivity, resolution, properties (web camera, action camera, and DSLR). RhythmEdge further stores longitudinal cardiovascular information and provides instant notification to the users. We thoroughly test the prototype stability, latency, and feasibility for three edge computing platforms by profiling their runtime, memory, and power usage.


Trends In Artificial Intelligence

#artificialintelligence

Artificial intelligence (AI) is a cutting-edge technology that is being adopted by forward-thinking businesses. The concept of artificial intelligence, on the other hand, has been around for decades. In 1955, "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence" was published, which coined the term "artificial intelligence." Dartmouth University sponsored the first AI research project in 1956, which is widely regarded as the start of artificial intelligence. So, why is AI gaining popularity now, more than sixty years later?


Artificial Intelligence Technology Trends That Matter for Business in 2022

#artificialintelligence

AI development has now reached a point where businesses of all sizes can use it, and it has a lot of promise. This blog talks about AI trends that companies can use and what experts think about the future of AI. People were curious about what was going to happen next in the field. There were a lot of expectations for AI after these jaw-dropping developments. This article will show you some of the most important AI developments that will make it more powerful and effective.



Applying AI to Real-Time Video Processing: The Basics and More

#artificialintelligence

If you look beyond image processing--it's one of the most common use cases for AI. And just like image processing, video processing uses established techniques like computer vision, object recognition, machine learning, and deep learning to enhance this process. Whether you use computer vision and NLP in video editing and generation, object recognition in video content auto-tagging tasks, machine learning to streamline AI video analysis, or deep learning to expedite real-time background removal, the use cases continue to grow by the day. Keep reading to learn what approach you can take when it comes to using AI in video processing. Let's start with the basics. Real-time video processing is an essential technology in surveillance systems using object and facial recognition.


'Liquid' machine-learning system adapts to changing conditions

#artificialintelligence

MIT researchers have developed a type of neural network that learns on the job, not just during its training phase. These flexible algorithms, dubbed "liquid" networks, change their underlying equations to continuously adapt to new data inputs. The advance could aid decision making based on data streams that change over time, including those involved in medical diagnosis and autonomous driving. "This is a way forward for the future of robot control, natural language processing, video processing--any form of time series data processing," says Ramin Hasani, the study's lead author. "The potential is really significant."


"Liquid" machine-learning system adapts to changing conditions

#artificialintelligence

MIT researchers have developed a type of neural network that learns on the job, not just during its training phase. These flexible algorithms, dubbed "liquid" networks, change their underlying equations to continuously adapt to new data inputs. The advance could aid decision making based on data streams that change over time, including those involved in medical diagnosis and autonomous driving. "This is a way forward for the future of robot control, natural language processing, video processing -- any form of time series data processing," says Ramin Hasani, the study's lead author. "The potential is really significant."


'Liquid' machine-learning system adapts to changing conditions: The new type of neural network could aid decision making in autonomous driving and medical diagnosis

#artificialintelligence

"This is a way forward for the future of robot control, natural language processing, video processing -- any form of time series data processing," says Ramin Hasani, the study's lead author. "The potential is really significant." The research will be presented at February's AAAI Conference on Artificial Intelligence. In addition to Hasani, a postdoc in the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL), MIT co-authors include Daniela Rus, CSAIL director and the Andrew and Erna Viterbi Professor of Electrical Engineering and Computer Science, and PhD student Alexander Amini. Other co-authors include Mathias Lechner of the Institute of Science and Technology Austria and Radu Grosu of the Vienna University of Technology.